Micron Document
<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Neurocomputational speech processing</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Neurocomputational_speech_processing"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Neurocomputational_speech_processing rootpage-Neurocomputational_speech_processing skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Neurocomputational speech processing</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr"><p><b>Neurocomputational speech processing</b> is computer-simulation of <a href="Speech_production" title="Speech production">speech production</a> and <a href="Speech_perception" title="Speech perception">speech perception</a> by referring to the natural neuronal processes of <a href="Speech_production" title="Speech production">speech production</a> and <a href="Speech_perception" title="Speech perception">speech perception</a>, as they occur in the human <a href="Nervous_system" title="Nervous system">nervous system</a> (<a href="Central_nervous_system" title="Central nervous system">central nervous system</a> and <a href="Peripheral_nervous_system" title="Peripheral nervous system">peripheral nervous system</a>). This topic is based on <a href="Neuroscience" title="Neuroscience">neuroscience</a> and <a href="Computational_neuroscience" title="Computational neuroscience">computational neuroscience</a>.<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Overview">Overview</h2></div>
<p>Neurocomputational models of speech processing are complex. They comprise at least a <a href="Cognition" title="Cognition">cognitive part</a>, a <a href="Motor_system" title="Motor system">motor part</a> and a <a href="Sensory_system" class="mw-redirect" title="Sensory system">sensory part</a>.<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p><p>The cognitive or linguistic part of a neurocomputational model of speech processing comprises the neural activation or generation of a <a href="Phonology" title="Phonology">phonemic representation</a> on the side of <a href="Speech_production" title="Speech production">speech production</a> (e.g. neurocomputational and extended version of the Levelt model developed by Ardi Roelofs:<sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> WEAVER++<sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> as well as the neural activation or generation of an intention or meaning on the side of <a href="Speech_perception" title="Speech perception">speech perception</a> or <a href="Reading_comprehension" title="Reading comprehension">speech comprehension</a>.
</p><p>The <a href="Motor_system" title="Motor system">motor part</a> of a neurocomputational model of speech processing starts with a <a href="Phonemic" class="mw-redirect" title="Phonemic">phonemic representation</a> of a speech item, activates a motor plan and ends with the <a href="Manner_of_articulation" title="Manner of articulation">articulation</a> of that particular speech item (see also: <a href="Articulatory_phonetics" title="Articulatory phonetics">articulatory phonetics</a>).
</p><p>The <a href="Sensory_system" class="mw-redirect" title="Sensory system">sensory part</a> of a neurocomputational model of speech processing starts with an acoustic signal of a speech item (<a href="Acoustic_phonetics" title="Acoustic phonetics">acoustic speech signal</a>), generates an <a href="Auditory_phonetics" title="Auditory phonetics">auditory representation</a> for that signal and activates a <a href="Phonemic" class="mw-redirect" title="Phonemic">phonemic representations</a> for that speech item.
</p>
<div class="mw-heading mw-heading2"><h2 id="Neurocomputational_speech_processing_topics">Neurocomputational speech processing topics</h2></div>
<p>Neurocomputational speech processing is speech processing by <a href="Artificial_neural_networks" class="mw-redirect" title="Artificial neural networks">artificial neural networks</a>. Neural maps, mappings and pathways as described below, are model structures, i.e. important structures within artificial neural networks.
</p>
<div class="mw-heading mw-heading3"><h3 id="Neural_maps">Neural maps</h3></div>

<p>An artificial neural network can be separated in three types of neural maps, also called "layers":
</p>
<ol><li>input maps (in the case of speech processing: primary auditory map within the <a href="Auditory_cortex" title="Auditory cortex">auditory cortex</a>, primary somatosensory map within the <a href="Somatosensory_cortex" class="mw-redirect" title="Somatosensory cortex">somatosensory cortex</a>),</li>
<li>output maps (primary motor map within the primary <a href="Motor_cortex" title="Motor cortex">motor cortex</a>), and</li>
<li>higher level cortical maps (also called "hidden layers").</li></ol>
<p>The term "neural map" is favoured here over the term "neural layer", because a cortical neural map should be modeled as a 2D-map of interconnected neurons (e.g. like a <a href="Self-organizing_map" title="Self-organizing map">self-organizing map</a>; see also Fig. 1). Thus, each "model neuron" or "<a href="Artificial_neuron" title="Artificial neuron">artificial neuron</a>" within this 2D-map is physiologically represented by a <a href="Cortical_column" title="Cortical column">cortical column</a> since the <a href="Cerebral_cortex" title="Cerebral cortex">cerebral cortex</a> anatomically exhibits a layered structure.
</p>
<div class="mw-heading mw-heading3"><h3 id="Neural_representations_(neural_states)">Neural representations (neural states)</h3></div>
<p>A neural representation within an <a href="Artificial_neural_network" class="mw-redirect" title="Artificial neural network">artificial neural network</a> is a temporarily activated (neural) state within a specific neural map. Each neural state is represented by a specific neural activation pattern. This activation pattern changes during speech processing (e.g. from syllable to syllable).
</p>

<p>In the ACT model (see below), it is assumed that an auditory state can be represented by a "neural <a href="Spectrogram" title="Spectrogram">spectrogram</a>" (see Fig. 2) within an auditory state map. This auditory state map is assumed to be located in the auditory association cortex (see <a href="Cerebral_cortex" title="Cerebral cortex">cerebral cortex</a>).
</p><p>A somatosensory state can be divided in a <a href="Touch" class="mw-redirect" title="Touch">tactile</a> and <a href="Proprioception" title="Proprioception">proprioceptive state</a> and can be represented by a specific neural activation pattern within the somatosensory state map. This state map is assumed to be located in the somatosensory association cortex (see <a href="Cerebral_cortex" title="Cerebral cortex">cerebral cortex</a>, <a href="Somatosensory_system" title="Somatosensory system">somatosensory system</a>, <a href="Somatosensory_cortex" class="mw-redirect" title="Somatosensory cortex">somatosensory cortex</a>).
</p><p>A motor plan state can be assumed for representing a motor plan, i.e. the planning of speech articulation for a specific syllable or for a longer speech item (e.g. word, short phrase). This state map is assumed to be located in the <a href="Premotor_cortex" title="Premotor cortex">premotor cortex</a>, while the instantaneous (or lower level) activation of each speech articulator occurs within the <a href="Primary_motor_cortex" title="Primary motor cortex">primary motor cortex</a> (see <a href="Motor_cortex" title="Motor cortex">motor cortex</a>).
</p><p>The neural representations occurring in the sensory and motor maps (as introduced above) are distributed representations (Hinton et al. 1968<sup id="cite_ref-5" class="reference"><a href="#cite_note-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>): Each neuron within the sensory or motor map is more or less activated, leading to a specific activation pattern.
</p><p>The neural representation for speech units occurring in the speech sound map (see below: DIVA model) is a punctual or local representation. Each speech item or speech unit is represented here by a specific <a href="Neuron" title="Neuron">neuron</a> (model cell, see below).
</p>
<div class="mw-heading mw-heading3"><h3 id="Neural_mappings_(synaptic_projections)">Neural mappings (synaptic projections)</h3></div>

<p>A neural mapping connects two cortical neural maps. Neural mappings (in contrast to neural pathways) store training information by adjusting their neural link weights (see <a href="Artificial_neuron" title="Artificial neuron">artificial neuron</a>, <a href="Artificial_neural_networks" class="mw-redirect" title="Artificial neural networks">artificial neural networks</a>). Neural mappings are capable of generating or activating a distributed representation (see above) of a sensory or motor state within a sensory or motor map from a punctual or local activation within the other map (see for example the synaptic projection from speech sound map to motor map, to auditory target region map, or to somatosensory target region map in the DIVA model, explained below; or see for example the neural mapping from phonetic map to auditory state map and motor plan state map in the ACT model, explained below and Fig. 3).
</p><p>Neural mapping between two neural maps are compact or dense: Each neuron of one neural map is interconnected with (nearly) each neuron of the other neural map (many-to-many-connection, see <a href="Artificial_neural_networks" class="mw-redirect" title="Artificial neural networks">artificial neural networks</a>). Because of this density criterion for neural mappings, neural maps which are interconnected by a neural mapping are not far apart from each other.
</p>
<div class="mw-heading mw-heading3"><h3 id="Neural_pathways">Neural pathways</h3></div>
<p>In contrast to neural mappings <a href="Neural_pathway" title="Neural pathway">neural pathways</a> can connect neural maps which are far apart (e.g. in different cortical lobes, see <a href="Cerebral_cortex" title="Cerebral cortex">cerebral cortex</a>). From the functional or modeling viewpoint, neural pathways mainly forward information without processing this information. A neural pathway in comparison to a neural mapping need much less neural connections. A neural pathway can be modelled by using a one-to-one connection of the neurons of both neural maps (see <a href="Topographic_map" title="Topographic map">topographic mapping</a> and see <a href="Somatotopic_arrangement" title="Somatotopic arrangement">somatotopic arrangement</a>).
</p><p>Example: In the case of two neural maps, each comprising 1,000 model neurons, a neural mapping needs up to 1,000,000 neural connections (many-to-many-connection), while only 1,000 connections are needed in the case of a neural pathway connection.
</p><p>Furthermore, the link weights of the connections within a neural mapping are adjusted during training, while the neural connections in the case of a neural pathway need not to be trained (each connection is maximal exhibitory).
</p>
<div class="mw-heading mw-heading2"><h2 id="DIVA_model">DIVA model</h2></div>
<p>The leading approach in neurocomputational modeling of speech production is the DIVA model developed by <a href="Frank_H._Guenther" title="Frank H. Guenther">Frank H. Guenther</a> and his group at Boston University.<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup> The model accounts for a wide range of <a href="Phonetic" class="mw-redirect" title="Phonetic">phonetic</a> and <a href="Neuroimaging" title="Neuroimaging">neuroimaging</a> data but - like each neurocomputational model - remains speculative to some extent.
</p>
<div class="mw-heading mw-heading3"><h3 id="Structure_of_the_model">Structure of the model</h3></div>
<p>The organization or structure of the DIVA model is shown in Fig. 4.
</p>
<div class="mw-heading mw-heading4"><h4 id="Speech_sound_map:_the_phonemic_representation_as_a_starting_point">Speech sound map: the phonemic representation as a starting point</h4></div>
<p>The speech sound map - assumed to be located in the inferior and posterior portion of <a href="Broca's_area" title="Broca's area">Broca's area</a> (left frontal operculum) - represents (phonologically specified) language-specific speech units (sounds, syllables, words, short phrases). Each speech unit (mainly syllables; e.g. the syllable and word "palm" /pam/, the syllables /pa/, /ta/, /ka/, ...) is represented by a specific model cell within the speech sound map (i.e. punctual neural representations, see above). Each model cell (see <a href="Artificial_neuron" title="Artificial neuron">artificial neuron</a>) corresponds to a small population of neurons which are located at close range and which fire together.
</p>
<div class="mw-heading mw-heading4"><h4 id="Feedforward_control:_activating_motor_representations">Feedforward control: activating motor representations</h4></div>
<p>Each neuron (model cell, <a href="Artificial_neuron" title="Artificial neuron">artificial neuron</a>) within the speech sound map can be activated and subsequently activates a forward motor command towards the motor map, called articulatory velocity and position map. The activated neural representation on the level of that motor map determines the articulation of a speech unit, i.e. controls all articulators (lips, tongue, velum, glottis) during the time interval for producing that speech unit. Forward control also involves subcortical structures like the <a href="Cerebellum" title="Cerebellum">cerebellum</a>, not modelled in detail here.
</p><p>A speech <i>unit</i> represents an amount of speech <i>items</i> which can be assigned to the same phonemic category. Thus, each speech unit is represented by one specific neuron within the speech sound map, while the realization of a speech unit may exhibit some articulatory and acoustic variability. This phonetic variability is the motivation to define sensory target <i>regions</i> in the DIVA model (see Guenther et al. 1998).<sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading4"><h4 id="Articulatory_model:_generating_somatosensory_and_auditory_feedback_information">Articulatory model: generating somatosensory and auditory feedback information</h4></div>
<p>The activation pattern within the motor map determines the movement pattern of all model articulators (lips, tongue, velum, glottis) for a speech item. In order not to overload the model, no detailed modeling of the <a href="Neuromuscular_junction" title="Neuromuscular junction">neuromuscular system</a> is done. The <a href="Articulatory_synthesis" title="Articulatory synthesis">Maeda articulatory speech synthesizer</a> is used in order to generate articulator movements, which allows the generation of a time-varying <a href="Vocal_tract" title="Vocal tract">vocal tract form</a> and the generation of the <a href="Acoustic_phonetics" title="Acoustic phonetics">acoustic speech signal</a> for each particular speech item.
</p><p>In terms of <a href="Artificial_intelligence" title="Artificial intelligence">artificial intelligence</a> the articulatory model can be called plant (i.e. the system, which is controlled by the brain); it represents a part of the <a href="Embodied_cognition" title="Embodied cognition">embodiment</a> of the neuronal speech processing system. The articulatory model generates <a href="Sensory_system" class="mw-redirect" title="Sensory system">sensory output</a> which is the basis for generating feedback information for the DIVA model (see below: feedback control).
</p>
<div class="mw-heading mw-heading4"><h4 id="Feedback_control:_sensory_target_regions,_state_maps,_and_error_maps">Feedback control: sensory target regions, state maps, and error maps</h4></div>
<p>On the one hand the articulatory model generates <a href="Sensory_system" class="mw-redirect" title="Sensory system">sensory information</a>, i.e. an auditory state for each speech unit which is neurally represented within the auditory state map (distributed representation), and a somatosensory state for each speech unit which is neurally represented within the somatosensory state map (distributed representation as well). The auditory state map is assumed to be located in the <a href="Temporal_cortex" class="mw-redirect" title="Temporal cortex">superior temporal cortex</a> while the somatosensory state map is assumed to be located in the <a href="Parietal_cortex" class="mw-redirect" title="Parietal cortex">inferior parietal cortex</a>.
</p><p>On the other hand, the speech sound map, if activated for a specific speech unit (single neuron activation; punctual activation), activates sensory information by synaptic projections between speech sound map and auditory target region map and between speech sound map and somatosensory target region map. Auditory and somatosensory target regions are assumed to be located in <a href="Auditory_cortex" title="Auditory cortex">higher-order auditory cortical regions</a> and in <a href="Somatosensory_cortex" class="mw-redirect" title="Somatosensory cortex">higher-order somatosensory cortical regions</a> respectively. These target region sensory activation patterns - which exist for each speech unit - are learned during <a href="Language_acquisition" title="Language acquisition">speech acquisition</a> (by imitation training; see below: learning).
</p><p>Consequently, two types of sensory information are available if a speech unit is activated at the level of the speech sound map: (i) learned sensory target regions (i.e. <i>intended</i> sensory state for a speech unit) and (ii) sensory state activation patterns resulting from a possibly imperfect execution (articulation) of a specific speech unit (i.e. <i>current</i> sensory state, reflecting the current production and articulation of that particular speech unit). Both types of sensory information is projected to sensory error maps, i.e. to an auditory error map which is assumed to be located in the <a href="Temporal_cortex" class="mw-redirect" title="Temporal cortex">superior temporal cortex</a> (like the auditory state map) and to a somatosensory error map which is assumed to be located in the <a href="Parietal_cortex" class="mw-redirect" title="Parietal cortex">inferior parietal cortex</a> (like the somatosensory state map) (see Fig. 4).
</p><p>If the current sensory state deviates from the intended sensory state, both error maps are generating feedback commands which are projected towards the motor map and which are capable to correct the motor activation pattern and subsequently the articulation of a speech unit under production. Thus, in total, the activation pattern of the motor map is not only influenced by a specific feedforward command learned for a speech unit (and generated by the synaptic projection from the speech sound map) but also by a feedback command generated at the level of the sensory error maps (see Fig. 4).
</p>
<div class="mw-heading mw-heading3"><h3 id="Learning_(modeling_speech_acquisition)">Learning (modeling speech acquisition)</h3></div>
<p>While the <i>structure</i> of a neuroscientific model of speech processing (given in Fig. 4 for the DIVA model) is mainly determined by <a href="Evolution" title="Evolution">evolutionary processes</a>, the (language-specific) <i>knowledge</i> as well as the (language-specific) <i>speaking skills</i> are learned and trained during <a href="Language_acquisition" title="Language acquisition">speech acquisition</a>. In the case of the DIVA model it is assumed that the newborn has not available an already structured (language-specific) speech sound map; i.e. no neuron within the speech sound map is related to any speech unit. Rather the organization of the speech sound map as well as the tuning of the projections to the motor map and to the sensory target region maps is learned or trained during speech acquisition. Two important phases of early speech acquisition are modeled in the DIVA approach: Learning by <a href="Babbling" title="Babbling">babbling</a> and by <a href="Imitation" title="Imitation">imitation</a>.
</p>
<div class="mw-heading mw-heading4"><h4 id="Babbling">Babbling</h4></div>
<p>During <a href="Babbling" title="Babbling">babbling</a> the synaptic projections between sensory error maps and motor map are tuned. This training is done by generating an amount of semi-random feedforward commands, i.e. the DIVA model "babbles". Each of these babbling commands leads to the production of an "articulatory item", also labeled as "pre-linguistic (i.e. non language-specific) speech item" (i.e. the articulatory model generates an articulatory movement pattern on the basis of the babbling motor command). Subsequently, an acoustic signal is generated.
</p><p>On the basis of the articulatory and acoustic signal, a specific auditory and somatosensory state pattern is activated at the level of the sensory state maps (see Fig. 4) for each (pre-linguistic) speech item. At this point the DIVA model has available the sensory and associated motor activation pattern for different speech items, which enables the model to tune the synaptic projections between sensory error maps and motor map. Thus, during babbling the DIVA model learns feedback commands (i.e. how to produce a proper (feedback) motor command for a specific sensory input).
</p>
<div class="mw-heading mw-heading4"><h4 id="Imitation">Imitation</h4></div>
<p>During <a href="Imitation" title="Imitation">imitation</a> the DIVA model organizes its speech sound map and tunes the synaptic projections between speech sound map and motor map - i.e. tuning of forward motor commands - as well as the synaptic projections between speech sound map and sensory target regions (see Fig. 4). Imitation training is done by exposing the model to an amount of acoustic speech signals representing realizations of language-specific speech units (e.g. isolated speech sounds, syllables, words, short phrases).
</p><p>The tuning of the synaptic projections between speech sound map and auditory target region map is accomplished by assigning one neuron of the speech sound map to the phonemic representation of that speech item and by associating it with the auditory representation of that speech item, which is activated at the auditory target region map. Auditory <i>regions</i> (i.e. a specification of the auditory variability of a speech unit) occur, because one specific speech unit (i.e. one specific phonemic representation) can be realized by several (slightly) different acoustic (auditory) realizations (for the difference between speech <i>item</i> and speech <i>unit</i> see above: feedforward control) .
</p><p>The tuning of the synaptic projections between speech sound map and motor map (i.e. tuning of forward motor commands) is accomplished with the aid of feedback commands, since the projections between sensory error maps and motor map were already tuned during babbling training (see above). Thus the DIVA model tries to "imitate" an auditory speech item by attempting to find a proper feedforward motor command. Subsequently, the model compares the resulting sensory output (<i>current</i> sensory state following the articulation of that attempt) with the already learned auditory target region (<i>intended</i> sensory state) for that speech item. Then the model updates the current feedforward motor command by the current feedback motor command generated from the auditory error map of the auditory feedback system. This process may be repeated several times (several attempts). The DIVA model is capable of producing the speech item with a decreasing auditory difference between current and intended auditory state from attempt to attempt.
</p><p>During imitation the DIVA model is also capable of tuning the synaptic projections from speech sound map to somatosensory target region map, since each new imitation attempt produces a new articulation of the speech item and thus produces a <a href="Somatosensory" class="mw-redirect" title="Somatosensory">somatosensory</a> state pattern which is associated with the phonemic representation of that speech item.
</p>
<div class="mw-heading mw-heading3"><h3 id="Perturbation_experiments">Perturbation experiments</h3></div>
<div class="mw-heading mw-heading4"><h4 id="Real-time_perturbation_of_F1:_the_influence_of_auditory_feedback">Real-time perturbation of F1: the influence of auditory feedback</h4></div>
<p>While auditory feedback is most important during speech acquisition, it may be activated less if the model has learned a proper feedforward motor command for each speech unit. But it has been shown that auditory feedback needs to be strongly coactivated in the case of auditory perturbation (e.g. shifting a formant frequency, Tourville et al. 2005).<sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup> This is comparable to the strong influence of visual feedback on reaching movements during visual perturbation (e.g. shifting the location of objects by viewing through a <a href="Prism_(optics)" title="Prism (optics)">prism</a>).
</p>
<div class="mw-heading mw-heading4"><h4 id="Unexpected_blocking_of_the_jaw:_the_influence_of_somatosensory_feedback">Unexpected blocking of the jaw: the influence of somatosensory feedback</h4></div>
<p>In a comparable way to auditory feedback, also somatosensory feedback can be strongly coactivated during speech production, e.g. in the case of unexpected blocking of the jaw (Tourville et al. 2005).
</p>
<div class="mw-heading mw-heading2"><h2 id="ACT_model">ACT model</h2></div>
<p>A further approach in neurocomputational modeling of speech processing is the ACT model developed by <a href="Bernd_J._Kr%C3%B6ger" title="Bernd J. Kröger">Bernd J. Kröger</a> and his group<sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup> at <a href="RWTH_Aachen_University" title="RWTH Aachen University">RWTH Aachen University</a>, Germany (Kröger et al. 2014,<sup id="cite_ref-13" class="reference"><a href="#cite_note-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup> Kröger et al. 2009,<sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup> Kröger et al. 2011<sup id="cite_ref-15" class="reference"><a href="#cite_note-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>). The ACT model is in accord with the DIVA model in large parts. The ACT model focuses on the "<a href="Motor_goal" title="Motor goal">action</a> repository" (i.e. <a href="Long-term_memory" title="Long-term memory">repository</a> for <a href="Motor_skills" class="mw-redirect" title="Motor skills">sensorimotor speaking skills</a>, comparable to the mental syllabary, see Levelt and Wheeldon 1994<sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup>), which is not spelled out in detail in the DIVA model. Moreover, the ACT model explicitly introduces a level of <a href="Motor_goal" title="Motor goal">motor plans</a>, i.e. a high-level motor description for the production of speech items (see <a href="Motor_goals" class="mw-redirect" title="Motor goals">motor goals</a>, <a href="Motor_cortex" title="Motor cortex">motor cortex</a>). The ACT model - like any neurocomputational model - remains speculative to some extent.
</p>
<div class="mw-heading mw-heading3"><h3 id="Structure">Structure</h3></div>

<p>The organization or structure of the ACT model is given in Fig. 5.
</p><p>For <a href="Speech_production" title="Speech production">speech production</a>, the ACT model starts with the activation of a <a href="Phonemic" class="mw-redirect" title="Phonemic">phonemic representation</a> of a speech item (phonemic map). In the case of a <i>frequent <a href="Syllable" title="Syllable">syllable</a></i>, a co-activation occurs at the level of the <a href="Phonetics" title="Phonetics">phonetic map</a>, leading to a further co-activation of the intended sensory state at the level of the <a href="Sensory_system" class="mw-redirect" title="Sensory system">sensory state maps</a> and to a co-activation of a <a href="Motor_system" title="Motor system">motor plan state</a> at the level of the motor plan map. In the case of an <i>infrequent syllable</i>, an attempt for a <a href="Motor_goal" title="Motor goal">motor plan</a> is generated by the motor planning module for that speech item by activating motor plans for phonetic similar speech items via the phonetic map (see Kröger et al. 2011<sup id="cite_ref-17" class="reference"><a href="#cite_note-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup>). The <a href="Motor_goal" title="Motor goal">motor plan</a> or vocal tract action score comprises temporally overlapping vocal tract actions, which are programmed and subsequently executed by the <a href="Motor_program" title="Motor program">motor programming, execution, and control module</a>. This module gets real-time somatosensory feedback information for controlling the correct execution of the (intended) motor plan. <a href="Motor_program" title="Motor program">Motor programing</a> leads to activation pattern at the level of the <a href="Primary_motor_cortex" title="Primary motor cortex">primary motor map</a> and subsequently activates <a href="Neuromuscular_junction" title="Neuromuscular junction">neuromuscular processing</a>. <a href="Motoneuron" class="mw-redirect" title="Motoneuron">Motoneuron activation patterns</a> generate <a href="Muscle" title="Muscle">muscle forces</a> and subsequently movement patterns of all <a href="Articulatory_phonetics" title="Articulatory phonetics">model articulators</a> (lips, tongue, velum, glottis). The <a href="Articulatory_synthesis" title="Articulatory synthesis">Birkholz 3D articulatory synthesizer</a> is used in order to generate the <a href="Acoustic_phonetics" title="Acoustic phonetics">acoustic speech signal</a>.
</p><p><a href="Articulatory_phonetics" title="Articulatory phonetics">Articulatory</a> and <a href="Acoustic_phonetics" title="Acoustic phonetics">acoustic</a> feedback signals are used for generating <a href="Somatosensory" class="mw-redirect" title="Somatosensory">somatosensory</a> and <a href="Auditory_system" title="Auditory system">auditory feedback information</a> via the sensory preprocessing modules, which is forwarded towards the auditory and somatosensory map. At the level of the sensory-phonetic processing modules, auditory and somatosensory information is stored in <a href="Short-term_memory" title="Short-term memory">short-term memory</a> and the external sensory signal (ES, Fig. 5, which are activated via the sensory feedback loop) can be compared with the already trained sensory signals (TS, Fig. 5, which are activated via the phonetic map). Auditory and somatosensory error signals can be generated if external and intended (trained) sensory signals are noticeably different (cf. DIVA model).
</p><p>The light green area in Fig. 5 indicates those neural maps and processing modules, which process a <a href="Syllable" title="Syllable">syllable</a> as a whole unit (specific processing time window around 100&nbsp;ms and more). This processing comprises the phonetic map and the directly connected sensory state maps within the sensory-phonetic processing modules and the directly connected motor plan state map, while the primary motor map as well as the (primary) auditory and (primary) somatosensory map process smaller time windows (around 10&nbsp;ms in the ACT model).
</p>

<p>The hypothetical <a href="Motor_cortex" title="Motor cortex">cortical location</a> of neural maps within the ACT model is shown in Fig. 6. The hypothetical locations of primary motor and primary sensory maps are given in magenta, the hypothetical locations of motor plan state map and sensory state maps (within sensory-phonetic processing module, comparable to the error maps in DIVA) are given in orange, and the hypothetical locations for the <a href="Mirror_neuron" title="Mirror neuron">mirrored</a> phonetic map is given in red. Double arrows indicate neuronal mappings. Neural mappings connect neural maps, which are not far apart from each other (see above). The two <a href="Mirror_neuron" title="Mirror neuron">mirrored</a> locations of the phonetic map are connected via a neural pathway (see above), leading to a (simple) one-to-one mirroring of the current activation pattern for both realizations of the phonetic map. This neural pathway between the two locations of the phonetic map is assumed to be a part of the <a href="Fasciculus_arcuatus" class="mw-redirect" title="Fasciculus arcuatus">fasciculus arcuatus</a> (AF, see Fig. 5 and Fig. 6).
</p><p>For <a href="Speech_perception" title="Speech perception">speech perception</a>, the model starts with an external acoustic signal (e.g. produced by an external speaker). This signal is preprocessed, passes the auditory map, and leads to an activation pattern for each syllable or word on the level of the auditory-phonetic processing module (ES: external signal, see Fig. 5). The ventral path of speech perception (see Hickok and Poeppel 2007<sup id="cite_ref-18" class="reference"><a href="#cite_note-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup>) would directly activate a lexical item, but is not implemented in ACT. Rather, in ACT the activation of a phonemic state occurs via the phonemic map and thus may lead to a coactivation of motor representations for that speech item (i.e. dorsal pathway of speech perception; ibid.).
</p>
<div class="mw-heading mw-heading4"><h4 id="Action_repository">Action repository</h4></div>

<p>The phonetic map together with the motor plan state map, sensory state maps (occurring within the sensory-phonetic processing modules), and phonemic (state) map form the action repository. The phonetic map is implemented in ACT as a <a href="Self-organizing_map" title="Self-organizing map">self-organizing neural map</a> and different speech items are represented by different neurons within this map (punctual or local representation, see above: neural representations). The phonetic map exhibits three major characteristics:
</p>
<ul><li>More than one <a href="Phonetics" title="Phonetics">phonetic realization</a> may occur within the phonetic map for one <a href="Phonemic" class="mw-redirect" title="Phonemic">phonemic state</a> (see phonemic link weights in Fig. 7: e.g. the syllable /de:m/ is represented by three neurons within the phonetic map)</li>
<li><a href="Phonetotopy" title="Phonetotopy">Phonetotopy</a>: The phonetic map exhibits an ordering of speech items with respect to different <a href="Phonetics" title="Phonetics">phonetic features</a> (see phonemic link weights in Fig. 7. Three examples: (i) the syllables /p@/, /t@/, and /k@/ occur in an upward ordering at the left side within the phonetic map; (ii) syllable-initial plosives occur in the upper left part of the phonetic map while syllable initial fricatives occur in the lower right half; (iii) CV syllables and CVC syllables as well occur in different areas of the phonetic map.).</li>
<li>The phonetic map is hypermodal or <a href="Multimodal_interaction" title="Multimodal interaction">multimodal</a>: The activation of a phonetic item at the level of the phonetic map coactivates (i) a phonemic state (see phonemic link weights in Fig. 7), (ii) a motor plan state (see motor plan link weights in Fig. 7), (iii) an auditory state (see auditory link weights in Fig. 7), and (iv) a somatosensory state (not shown in Fig. 7). All these states are learned or trained during speech acquisition by tuning the synaptic link weights between each neuron within the phonetic map, representing a particular phonetic state and all neurons within the associated motor plan and sensory state maps (see also Fig. 3).</li></ul>
<p>The phonetic map implements the <a href="Action-Specific_Perception" class="mw-redirect" title="Action-Specific Perception">action-perception-link</a> within the ACT model (see also Fig. 5 and Fig. 6: the dual neural representation of the phonetic map in the <a href="Frontal_lobe" title="Frontal lobe">frontal lobe</a> and at the intersection of <a href="Temporal_lobe" title="Temporal lobe">temporal lobe</a> and <a href="Parietal_lobe" title="Parietal lobe">parietal lobe</a>).
</p>
<div class="mw-heading mw-heading4"><h4 id="Motor_plans">Motor plans</h4></div>
<p>A motor plan is a high level motor description for the production and articulation of a speech items (see <a href="Motor_goals" class="mw-redirect" title="Motor goals">motor goals</a>, <a href="Motor_skills" class="mw-redirect" title="Motor skills">motor skills</a>, <a href="Articulatory_phonetics" title="Articulatory phonetics">articulatory phonetics</a>, <a href="Articulatory_phonology" title="Articulatory phonology">articulatory phonology</a>). In our neurocomputational model ACT a motor plan is quantified as a vocal tract action score. Vocal tract action scores quantitatively determine the number of vocal tract actions (also called articulatory gestures), which need to be activated in order to produce a speech item, their degree of realization and duration, and the temporal organization of all vocal tract actions building up a speech item (for a detailed description of vocal tract actions scores see e.g. Kröger &amp; Birkholz 2007).<sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup> The detailed realization of each vocal tract action (articulatory gesture) depends on the temporal organization of all vocal tract actions building up a speech item and especially on their temporal overlap. Thus the detailed realization of each vocal tract action within a speech item is specified below the motor plan level in our neurocomputational model ACT (see Kröger et al. 2011).<sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Integrating_sensorimotor_and_cognitive_aspects:_the_coupling_of_action_repository_and_mental_lexicon">Integrating sensorimotor and cognitive aspects: the coupling of action repository and mental lexicon</h3></div>
<p>A severe problem of phonetic or sensorimotor models of speech processing (like DIVA or ACT) is that the development of the <a href="Phonemic" class="mw-redirect" title="Phonemic">phonemic map</a> during speech acquisition is not modeled. A possible solution of this problem could be a direct coupling of action repository and mental lexicon without explicitly introducing a phonemic map at the beginning of speech acquisition (even at the beginning of imitation training; see Kröger et al. 2011 PALADYN Journal of Behavioral Robotics).
</p>
<div class="mw-heading mw-heading3"><h3 id="Experiments:_speech_acquisition">Experiments: speech acquisition</h3></div>
<p>A very important issue for all neuroscientific or neurocomputational approaches is to separate structure and knowledge. While the structure of the model (i.e. of the human neuronal network, which is needed for processing speech) is mainly determined by <a href="Evolution" title="Evolution">evolutionary processes</a>, the knowledge is gathered mainly during <a href="Language_acquisition" title="Language acquisition">speech acquisition</a> by processes of <a href="Learning" title="Learning">learning</a>. Different learning experiments were carried out with the model ACT in order to learn (i) a five-vowel system /i, e, a, o, u/ (see Kröger et al. 2009), (ii) a small consonant system (voiced plosives /b, d, g/ in combination with all five vowels acquired earlier as CV syllables (ibid.), (iii) a small model language comprising the five-vowel system, voiced and unvoiced plosives /b, d, g, p, t, k/, nasals /m, n/ and the lateral /l/ and three syllable types (V, CV, and CCV) (see Kröger et al. 2011)<sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup> and (iv) the 200 most frequent syllables of Standard German for a 6-year-old child (see Kröger et al. 2011).<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup> In all cases, an ordering of phonetic items with respect to different phonetic features can be observed.
</p>
<div class="mw-heading mw-heading3"><h3 id="Experiments:_speech_perception">Experiments: speech perception</h3></div>
<p>Despite the fact that the ACT model in its earlier versions was designed as a pure speech production model (including speech acquisition), the model is capable of exhibiting important basic phenomena of speech perception, i.e. categorical perception and the McGurk effect. In the case of <a href="Categorical_perception" title="Categorical perception">categorical perception</a>, the model is able to exhibit that categorical perception is stronger in the case of plosives than in the case of vowels (see Kröger et al. 2009). Furthermore, the model ACT was able to exhibit the <a href="McGurk_effect" title="McGurk effect">McGurk effect</a>, if a specific mechanism of inhibition of neurons of the level of the phonetic map was implemented (see Kröger and Kannampuzha 2008).<sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1290876196">
/* start https://en.wikipedia.org/ */


.mw-parser-output .side-box{margin:4px 0;box-sizing:border-box;border:1px solid #aaa;font-size:88%;line-height:1.25em;background-color:var(--background-color-interactive-subtle,#f8f9fa);display:flow-root}.mw-parser-output .infobox .side-box{font-size:100%}.mw-parser-output .side-box-abovebelow,.mw-parser-output .side-box-text{padding:0.25em 0.9em}.mw-parser-output .side-box-image{padding:2px 0 2px 0.9em;text-align:center}.mw-parser-output .side-box-imageright{padding:2px 0.9em 2px 0;text-align:center}@media(min-width:500px){.mw-parser-output .side-box-flex{display:flex;align-items:center}.mw-parser-output .side-box-text{flex:1;min-width:0}}@media(min-width:720px){.mw-parser-output .side-box{width:238px}.mw-parser-output .side-box-right{clear:right;float:right;margin-left:1em}.mw-parser-output .side-box-left{margin-right:1em}}


/* end https://en.wikipedia.org/ */
</style><style data-mw-deduplicate="TemplateStyles:r1237033735">
/* start https://en.wikipedia.org/ */


@media print{body.ns-0 .mw-parser-output .sistersitebox{display:none!important}}@media screen{html.skin-theme-clientpref-night .mw-parser-output .sistersitebox img[src*="Wiktionary-logo-en-v2.svg"]{background-color:white}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .sistersitebox img[src*="Wiktionary-logo-en-v2.svg"]{background-color:white}}


/* end https://en.wikipedia.org/ */
</style><div class="side-box side-box-right sistersitebox"><style data-mw-deduplicate="TemplateStyles:r1126788409">
/* start https://en.wikipedia.org/ */


.mw-parser-output .plainlist ol,.mw-parser-output .plainlist ul{line-height:inherit;list-style:none;margin:0;padding:0}.mw-parser-output .plainlist ol li,.mw-parser-output .plainlist ul li{margin-bottom:0}


/* end https://en.wikipedia.org/ */
</style>
<div class="side-box-flex">
<div class="side-box-image"><span class="noviewer" typeof="mw:File"></span></div>
<div class="side-box-text plainlist">Wikimedia Commons has media related to <span style="font-weight: bold; font-style: italic;"><a href="https://commons.wikimedia.org/wiki/Category:Neurocomputational_speech_processing" class="extiw external" title="commons:Category:Neurocomputational speech processing">Neurocomputational speech processing</a></span>.</div></div>
</div>
<ul><li><a href="Speech_production" title="Speech production">Speech production</a></li>
<li><a href="Speech_perception" title="Speech perception">Speech perception</a></li>
<li><a href="Computational_neuroscience" title="Computational neuroscience">Computational neuroscience</a></li>
<li><a href="Articulatory_synthesis" title="Articulatory synthesis">Articulatory synthesis</a></li>
<li><a href="Auditory_feedback" title="Auditory feedback">Auditory feedback</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */


.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}


/* end https://en.wikipedia.org/ */
</style><div class="reflist reflist-columns references-column-width reflist-columns-2">
<ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */


.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}


/* end https://en.wikipedia.org/ */
</style><cite class="citation book cs1"><a rel="nofollow" class="external text" href="https://dl.acm.org/doi/10.5555/1768226.1768230">"Towards neurocomputational speech and sound processing"</a>. <i>Progress in nonlinear speech processing</i>. Springer. January 2007. pp.&nbsp;<span class="nowrap">58–</span>77. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-3-540-71503-0</bdi>.</cite></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><cite id="CITEREFParrellLammertCiccarelliQuatieri2019" class="citation journal cs1">Parrell, Benjamin; Lammert, Adam C.; Ciccarelli, Gregory; Quatieri, Thomas F. (2019-03-01). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://doi.org/10.1121/1.5092807">"Current models of speech motor control: A control-theoretic overview of architectures and properties"</a></span>. <i>The Journal of the Acoustical Society of America</i>. <b>145</b> (3): <span class="nowrap">1456–</span>1481. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2019ASAJ..145.1456P">2019ASAJ..145.1456P</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1121%2F1.5092807">10.1121/1.5092807</a>. <a href="ISSN_(identifier)" class="mw-redirect" title="ISSN (identifier)">ISSN</a>&nbsp;<a rel="nofollow" class="external text" href="https://search.worldcat.org/issn/0001-4966">0001-4966</a>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/31067944">31067944</a>.</cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://web.archive.org/web/20120426040535/http://www.nici.kun.nl/~ardiroel/home.htm">"Ardi Roelofs"</a>. Archived from <a rel="nofollow" class="external text" href="http://www.nici.kun.nl/~ardiroel/home.htm">the original</a> on 2012-04-26<span class="reference-accessdate">. Retrieved <span class="nowrap">2011-12-08</span></span>.</cite></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-4">^</a></b></span> <span class="reference-text"><a rel="nofollow" class="external text" href="http://www.nici.ru.nl/~ardiroel/weaver++.htm">WEAVER++</a></span>
</li>
<li id="cite_note-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-5">^</a></b></span> <span class="reference-text">Hinton GE, McClelland JL, Rumelhart DE (1968) Distributed representations. In: Rumelhart DE, McClelland JL (eds.). <i>Parallel Distributed Processing: Explorations in the Microstructure of Cognition</i>. Volume 1: Foundations (MIT Press, Cambridge, MA)</span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-6">^</a></b></span> <span class="reference-text">DIVA model: a model of speech production, focussing on feedback control processes, developed by <a rel="nofollow" class="external text" href="http://cns.bu.edu/~guenther/">Frank H. Guenther and his group at Boston University, MA, USA</a>. The term "DIVA" refers to "Directions Into Velocities of Articulators"</span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text">Guenther, F.H., Ghosh, S.S., and Tourville, J.A. (2006) <a rel="nofollow" class="external text" href="http://cns.bu.edu/~guenther/Guenther_et_al_Brain-Lang_2005.pdf">pdf</a> <a rel="nofollow" class="external text" href="https://web.archive.org/web/20120415055307/http://cns.bu.edu/~guenther/Guenther_et_al_Brain-Lang_2005.pdf">Archived</a> 2012-04-15 at the <a href="Wayback_Machine" title="Wayback Machine">Wayback Machine</a>. Neural modeling and imaging of the cortical interactions underlying syllable production. <i>Brain and Language</i>, 96, pp.&nbsp;280–301</span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-8">^</a></b></span> <span class="reference-text">Guenther FH (2006) Cortical interaction underlying the production of speech sounds. <i>Journal of Communication Disorders</i> 39, 350–365</span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text">Guenther, F.H., and Perkell, J.S. (2004) <a rel="nofollow" class="external text" href="http://cns.bu.edu/~guenther/SMC_2001_Book_Chapter.pdf">pdf</a> <a rel="nofollow" class="external text" href="https://web.archive.org/web/20120415055319/http://cns.bu.edu/~guenther/SMC_2001_Book_Chapter.pdf">Archived</a> 2012-04-15 at the <a href="Wayback_Machine" title="Wayback Machine">Wayback Machine</a>. A neural model of speech production and its application to studies of the role of auditory feedback in speech. In: B. Maassen, R. Kent, H. Peters, P. Van Lieshout, and W. Hulstijn (eds.), <i>Speech Motor Control in Normal and Disordered Speech</i> (pp.&nbsp;29–49). Oxford: Oxford University Press</span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><b><a href="#cite_ref-10">^</a></b></span> <span class="reference-text"><cite id="CITEREFGuentherHampsonJohnson1998" class="citation journal cs1">Guenther, Frank H.; Hampson, Michelle; Johnson, Dave (1998). "A theoretical investigation of reference frames for the planning of speech movements". <i>Psychological Review</i>. <b>105</b> (4): <span class="nowrap">611–</span>633. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1037%2F0033-295x.105.4.611-633">10.1037/0033-295x.105.4.611-633</a>. <a href="Hdl_(identifier)" class="mw-redirect" title="Hdl (identifier)">hdl</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://hdl.handle.net/2144%2F2114">2144/2114</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/9830375">9830375</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:11179837">11179837</a>.</cite></span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-11">^</a></b></span> <span class="reference-text">Tourville J, Guenther F, Ghosh S, Reilly K, Bohland J, Nieto-Castanon A (2005) Effects of acoustic and articulatory perturbation on cortical activity during speech production. <i>Poster, 11th annual meeting of the Organization of Human Brain Mapping</i> (Toronto, Canada)</span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-12">^</a></b></span> <span class="reference-text">ACT model: A model of speech production, perception, and acquisition, developed by <a rel="nofollow" class="external text" href="http://www.phonetik.phoniatrie.rwth-aachen.de/bkroeger/">Bernd J. Kröger and his group at RWTH Aachen University, Germany</a>. The term "ACT" refers to the term "ACTion"</span>
</li>
<li id="cite_note-13"><span class="mw-cite-backlink"><b><a href="#cite_ref-13">^</a></b></span> <span class="reference-text">BJ Kröger, J Kannampuzha, E Kaufmann (2014) <a rel="nofollow" class="external text" href="http://www.phonetik.phoniatrie.rwth-aachen.de/bkroeger/documents/Kroeger_2014_EPJNBP.pdf">pdf</a> Associative learning and self-organization as basic principles for simulating speech acquisition, speech production, and speech perception. EPJ Nonlinear Biomedical Physics 2 (1), 1-28</span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-14">^</a></b></span> <span class="reference-text">Kröger BJ, Kannampuzha J, Neuschaefer-Rube C (2009) <a rel="nofollow" class="external text" href="http://www.phonetik.phoniatrie.rwth-aachen.de/bkroeger/documents/Kroeger_etal_2009.pdf">pdf</a> Towards a neurocomputational model of speech production and perception. <i>Speech Communication</i> 51: 793-809</span>
</li>
<li id="cite_note-15"><span class="mw-cite-backlink"><b><a href="#cite_ref-15">^</a></b></span> <span class="reference-text"><cite id="CITEREFKrögerBirkholzNeuschaefer-Rube2011" class="citation journal cs1">Kröger, Bernd J.; Birkholz, Peter; Neuschaefer-Rube, Christiane (1 June 2011). <a rel="nofollow" class="external text" href="https://doi.org/10.2478%2Fs13230-011-0016-6">"Towards an Articulation-Based Developmental Robotics Approach for Word Processing in Face-to-Face Communication"</a>. <i>Paladyn. Journal of Behavioral Robotics</i>. <b>2</b> (2): <span class="nowrap">82–</span>93. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.2478%2Fs13230-011-0016-6">10.2478/s13230-011-0016-6</a></span>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:10317127">10317127</a>.</cite></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><b><a href="#cite_ref-16">^</a></b></span> <span class="reference-text"><cite id="CITEREFLeveltWheeldon1994" class="citation journal cs1">Levelt, Willem J.M.; Wheeldon, Linda (April 1994). "Do speakers have access to a mental syllabary?". <i>Cognition</i>. <b>50</b> (<span class="nowrap">1–</span>3): <span class="nowrap">239–</span>269. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2F0010-0277%2894%2990030-2">10.1016/0010-0277(94)90030-2</a>. <a href="Hdl_(identifier)" class="mw-redirect" title="Hdl (identifier)">hdl</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://hdl.handle.net/2066%2F15533">2066/15533</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/8039363">8039363</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:7845880">7845880</a>.</cite></span>
</li>
<li id="cite_note-17"><span class="mw-cite-backlink"><b><a href="#cite_ref-17">^</a></b></span> <span class="reference-text">Kröger BJ, Miller N, Lowit A, Neuschaefer-Rube C. (2011) Defective neural motor speech mappings as a source for apraxia of speech: Evidence from a quantitative neural model of speech processing. In: Lowit A, Kent R (eds.) Assessment of Motor Speech Disorders. (Plural Publishing, San Diego, CA) pp. 325-346</span>
</li>
<li id="cite_note-18"><span class="mw-cite-backlink"><b><a href="#cite_ref-18">^</a></b></span> <span class="reference-text">Hickok G, Poeppel D (2007) Towards a functional neuroanatomy of speech perception. <i>Trends in Cognitive Sciences</i> 4, 131–138</span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><b><a href="#cite_ref-19">^</a></b></span> <span class="reference-text">Kröger BJ, Birkholz P (2007) A gesture-based concept for speech movement control in articulatory speech synthesis. In: Esposito A, Faundez-Zanuy M, Keller E, Marinaro M (eds.) <i>Verbal and Nonverbal Communication Behaviours, LNAI 4775</i> (Springer Verlag, Berlin, Heidelberg) pp. 174-189</span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><b><a href="#cite_ref-20">^</a></b></span> <span class="reference-text">Kröger BJ, Birkholz P, Kannampuzha J, Eckers C, Kaufmann E, Neuschaefer-Rube C (2011) Neurobiological interpretation of a quantitative target approximation model for speech actions. In: Kröger BJ, Birkholz P (eds.) <i>Studientexte zur Sprachkommunikation: Elektronische Sprachsignalverarbeitung 2011</i> (TUDpress, Dresden, Germany), pp. 184-194</span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><b><a href="#cite_ref-21">^</a></b></span> <span class="reference-text">Kröger BJ, Miller N, Lowit A, Neuschaefer-Rube C. (2011) Defective neural motor speech mappings as a source for apraxia of speech: Evidence from a quantitative neural model of speech processing. In: Lowit A, Kent R (eds.) <i>Assessment of Motor Speech Disorders.</i> (Plural Publishing, San Diego, CA) pp. 325-346</span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><b><a href="#cite_ref-22">^</a></b></span> <span class="reference-text">Kröger BJ, Birkholz P, Kannampuzha J, Kaufmann E, Neuschaefer-Rube C (2011) Towards the acquisition of a sensorimotor vocal tract action repository within a neural model of speech processing. In: Esposito A, Vinciarelli A, Vicsi K, <a href="Catherine_Pelachaud" title="Catherine Pelachaud">Pelachaud C</a>, Nijholt A (eds.) <i>Analysis of Verbal and Nonverbal Communication and Enactment: The Processing Issues.</i> LNCS 6800 (Springer, Berlin), pp. 287-293</span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><b><a href="#cite_ref-23">^</a></b></span> <span class="reference-text">Kröger BJ, Kannampuzha J (2008) A neurofunctional model of speech production including aspects of auditory and audio-visual speech perception. <i>Proceedings of the International Conference on Audio-Visual Speech Processing 2008</i> (Moreton Island, Queensland, Australia) pp. 83–88</span>
</li>
</ol></div>
<div class="mw-heading mw-heading2"><h2 id="Further_reading">Further reading</h2></div>
<ul><li><a rel="nofollow" class="external text" href="https://www.researchgate.net/profile/Iaroslav_Blagouchine/publication/224080014_Control_of_a_Speech_Robot_via_an_Optimum_Neural-Network-Based_Internal_Model_With_Constraints">Iaroslav Blagouchine and Eric Moreau. <i>Control of a Speech Robot via an Optimum Neural-Network-Based Internal Model with Constraints.</i> IEEE Transactions on Robotics, vol. 26, no. 1, pp. 142—159, February 2010.</a></li></ul></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-07-21" href="https://en.wikipedia.org/wiki/?title=Neurocomputational_speech_processing&amp;oldid=1301699403">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>

</body></html>